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(57) A processor-based device (1 02) incorporating 
an on-chip instruction trace cache (200) capable of pro- 
viding Infonnatioh for reconstructing Instruction execu- 
tion flow. The trace information can be captured without 
halting normal processor (104) operation. Both serial 
(204) and parallel (214) communication channels are 
provided for communicating the trace Infonnation to ex- 
ternal devices. In the disclosed embodiment of the In- 
vention, instructions that disrupt the instruction flow are 
reported, particularly instructions in which the target ad- 
dress is in some way data dependent. For example, call 



instructions or unconditional branch Instructions In 
which tfie target address is provided from a data register 
(or othermemory location such as a stack) cause atrace 
cache entry to be generated. In tiie case of many un- 
conditional branches or sequential instnjctions, no entiy 
is placed into the trace cache (200) because the target 
address can be completely detenmlned from the Instruc- 
tion stream. Other Information provided by the Instruc- 
tion trace cache (200) includes: the target address of a 
trap or Intenrupt handler, the target address of a return 
instruction, addresses from procedure returns, task 
identifiers, and trace capture stop/start infonnation. 
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EP1 184 790A2 

Description 
TECHNICAL FIELD 

5 [0001 ] The Invention relates to software debug support In microprocessors, and more particularly to a microproces- 
sor-based device incorporating an on^hip instruction trace cache. 

BACKGROUND AFTT 

10 [0002] The growth In software complexity, coupled with Increasing processor clock speeds, has placed an Increasing 
burden on application software developers. The cost of developing and debugging new software products is now a 
significant factor in processor selection. A processor's failure to adequately facilitate software debug results in longer 
customer development times and reduces the processor's attractiveness for use within industry. The need to provide 
software debug support is particularly acute within the embedded products industry, where specialized on-chip circuitry 

15 is often combined with a p'rocessbr'core. 

[0003] In addition to the software engineer, other parties are also affected by debug tool configuration. These parties 
include: the "trace" algorithm developer who must search through captured software trace data that reflects instruction 
execution flow In a processor; the in-clrcult emulator developer who deals with problems of signal synchronization, 
clocic frequency and trace bandwidth; and the processor manufacturer who does not want a solution that results In 

20 increased processor cost or design and development complexity. 

[0004] With desktop systems, complex multitasking operating systems are currently available to support debugging. 
However, the initial task of getting these operating systems running reliably often requires special development equip- 
ment. While not the standard in the desktop environment, the use of such equipment is often the approach taken within 
the embedded industry. Logic analyzers, read-only memory (ROM) emulators and In-circuft emulators (ICE) are fre- 

25 q uently employed. In-clrcuit emulators do provide certain advantages over other debug environments, offering complete 
control and visibility over memory and register contents, as well as overlay and trace memory in case system memory 
Is insufficient Use of traditional in-circuit emulators, which involves Interfacing a custom emulator back-end with a 
processor socket to allow communication between emulation equipment and the target system, is becoming increas- 
ingly difficult and expensive In today's age of exotic packages and shrinking product life cycles. 

30 [0005] Assuming full-function In-circult emulation is required, there are a few known processor manufacturing tech- 
niques able to offer the required support for emulation equipment. Most processors Intended for personal computer 
(PC) systems utilize a multiplexed approach in which existing pins are multiplexed for use in software debug. This 
approach is not particularly desirable in the embedded industry, where It Is more difficult to overioad pin functionality. 
[0006] Other more advanced processors multiplex debug pins In time. In such processors, the address bus Is used 

35 to report software trace Infomnation during a BTA-cycle (Branch Target Address). The BTA-cycle, however, must be 
stolen from the regular bus operation. In debug environments where branch activity is high and cache hit rates are low, 
it becomes impossible to hide the BTA-cycles. The resulting confitet over access to the address bus necessitates 
processor "throttle back" to prevent loss of instruction trace infomnation. In the communteations industry, for example, 
software typically makes extensive use of branching and suffers poor cache utilization, often resulting In 20% throttle 

40 back or more. This amount of throttling is unacceptable amount for embedded products which must accommodate 
reai-time constrains. 

[0007] In another approach, a second "trace" or "slave" processor Is combined with the main processor, with the two 
processors operating In-step. Only the main processor is required to fetch Instnictlons. The second, slave processor 
is used to monitor the fetched instmctlons on the data bus and keeps Its internal state In synchronization with the main 

^5 processor. The address bus of the slave processor functions to provide trace Infonnatlon. After power-up, via a JTAG 
(Joint Test Action Group) input, the second processor is switched into a slave mode of operation. Free from the need 
to fetch Instnjctions, its address bus and other pins provide the necessary trace Infonmation. 
[0008] Another existing approach involves building debug support Into every processor, but only bonding-out the 
necessary signal pins in a limited number of packages. These "specially" packaged versions of the processor are used 

so during debug and replaced with the smaller package for final production. This bond-out approach suffers from the need 
to support additional bond pad sites in all fabricated devices. This can be a burden in small packages and pad limited 
designs, particularly if a substantial number of "extra" pins are required by the debug support variant. Additionally, the 
debug capability ofthe specially packaged processors Is unavailable In typical processor-based production systems. 
[0009] In yet another approach (the "Background Debug Mode" by Motorola, Inc.) limited on-chip debug circuitry Is 

55 provided for basic run control. Through a dedicated serial link requiring additional pins, this approach allows a debugger 
to start and stop the target system and apply bask: code breakpoints by inserting special instructions In system memory. 
Once halted, special commands are used to inspect memory variables and register contents. This serial link, however,, 
does not provide trace support - additional dedicated pins and expensive external trace capture hardware are required 
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to provide instruction trace data. European patent application EP-A-0 762 276 of Motorola describes a debug module 
of a data processor which provides a parallel output port for providing internal operating Information via a DDATA signal 
and a PST signal. The DDATA signal provides data which reflects operand values and the PST signal provides encoded 
status Information which reflects an execution status of the central processing unit. 
s pOlO] Thus, the current solutions for software debugging suffer from a variety of limitations, Including: Increased . 
packaging and development costs, circuit complexity, processorthrottling, and bandwidth matching difficulties. Further, 
there Is cun^ntiy no adequate low-cost procedure for providing trace infomiation. The limitations of the existing solutions 
are likely to be exacerbated In the future as Internal processor clock frequencies continue to Increase. 

10 DISCLOSURE OF THE INVENTION 

[0011] Briefly, a processor-based device according to the present Invention includes an on-chip instruction trace 
cache capable of providing Infonnatlon for reconstructing Instruction execution flow. The trace information can be cap- 
tured without halting normal processor operation. Both serial and parallel communication channels are provided for 
IS communicating the trace Information to extemat devices. In tHe discl osed e mbodiment of the Invention, controHabHity 
and observability of the instruction trace cache are achieved through a software debug port that uses an IEEE- 1 
149.1-1990 compliant JTAG (Joint Test Action Group) interface or a similar standardized interface that is Integrated 
Into the processor-based device. 

[0012] According to a first aspect, the present invention provides an electronic processor-based device adapted to 
20 execute a series of Instructions obtained from external sources, the processor-based device being provided with pins 
to pennit connection to extemal conductors, the electron^ processor-based device being characterized by: 

a trace cache coupled to a processor core for storing trace infomnatlon indicative of the order in which the instruc- 
tions are executed by the processor core, the trace cache comprising a series of storage elements, each storage 
2s element being adapted to store trace Infonnatlon, the trace Information Including a plurality of instruction trace 
records containing address and data Infomriatfon, the trace cache being configured to load data from the processor 
core In response to a load command and configured for the processor core to retrieve data from the trace cache 
In response to a retrieve command; 

and a communication Interface connected between the trace cache and selected ones of the pins to provide for 
30 transmission of trace Infonnatlon from the trace cache to extemal devices. 

[0013] According to a second aspect the present invention provides method for analysing trace Infonnation in a 
processor-based device having a processor core comprising the steps of: 

35 providing a trace cache within the processor-based device the trace cache comprising a series of storage elements 
adapted to store trace Infonnation ; 

capturing trace Infonnation from the processor core that is indicative of the order in which the series of Instructions 
is executed by the processor core; 

storing the trace infonnation in the trace cache storage elements as instruction trace records; 
40 retrieving the trace Infomiation by the processor core from the trace cache storage elements in response to a 
retrieve command; and 

loading other Infomiation from the processor core Into the trace cache storage elements in response to a load 
command; 

providing a communication channel from the trace cache to selected pins of the processor-based device and 
45 communicating the trace infomiation from the trace cache to the selected pins via tho communication channel. 

[0014] Preferably, infonnatlon stored In the Instruction trace cache is "compressed" such that a smaller cache can 
be utilized. In addition, compressing trace data allows external hardware to operate at nonmai bus speeds, even while 
the internal processor is operating much faster. Less expensive external capture hardware can therefore be utilized 

so with a processor-based device according to the Invention. 

[0015] In the disclosed embodiment of the Invention, If an address In an instruction stream can be obtained from a 
program image (Object l^odule), then it is not provided in the trace data. Preferably, only instructions that disrupt the 
Instnjction flow are reported; and further, only instajctions in which the target address is In some way data dependent. 
Such "dismpting" events include, for example, call instructions or unconditional branch instructions In which the target 

55 address Is provided from a data register or other memory location such as a stack. In the case of many unconditional 
branches or sequential Instmctlons, no entry Is placed into the trace cache because the target address can be com- 
pletely determined from the Instmction stream. Other Infonmatlon provided by the instruction trace cache includes: the 
target address of a trap or Intenupt handler, the target address of a return instruction, addresses from procedure returns, 
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task Identifiers, and trace capture stop/start Information. This technique reduces the amount of Infonnatlon transfen'ed 
from the trace cache to external debug hardware. 

[001 6] Thus, a processor-based device supplying a flexible, high-perfomiance solution for furnishing Instnjctlon trace 
Infomiatlon is provided by the Invention. The disclosed on-chip Instruction trace cache alleviates various of the band- 
5 width and clock synchronization problems that arise in many existing solutions. 

BRIEF DESCRIPTION OF DRAWINGS 

[0017] A better understanding of the present Invention can be obtained when the following detailed description of 
10 the prefen'ed embodiment is considered in conjunction with the following drawings, in which: 

Figure 1 Is a block diagram of a software debug environment utilizing a software debug solution In accordance 
with the present invention; 

Figure 2 Is a block diagram providing details of an exemplary embedded processor product Incorporating an on* 
f 5 chlp'fnstnictlon trace cache according to thepresent inventfon; 

Rgure 3 is a simplified block diagram depicting the relationship between an exemplary Instruction trace cache and 

other components of an embedded processor product according to the present invention; 

Rgure 4 is a flowchart Illustrating software debug command passing according to one embodiment of the invention; 

Rgure 5 Is a flowchart Illustrating enhanced software port command passing according to a second embodiment 
20 of the invention; and 

Rgures 6A - 6G illustrate the general format of a variety of trace cache entries for reporting Instnjc^lon execution 

according to the Invention. 

MODE(S) FOR CARRYING OUT THE INVE^r^ON 

25 

[0018] Turning now to the drawings. Figure 1 depicts an exemplary software debug environment Illustrating a con- 
templated use of the present Invention. A target system T is shown containing an embedded processor device 102 
according to the present invention coupled to system memory 1 06. The embedded processor device 1 02 Incorporates 
a processor core 1 04, an Instruction trace cache 200 (Rgure 2), and a debug port 1 00. Although not considered critical 
30 to the invention , the embedded processor device 1 02 may Incorporate additional circuitry (not shown) for performing 
application specific functions, or may take the fonri of a stand-alone processor or digital signal processor. Preferably, 
the debug port 100 uses an IEEE-1149.1-1990 compliant JTAG Interface or other similar standardized serial port In- 
terface. 

[0019] A host system H Is used to execute debug control software 112 for transferring high-level commands and 

3S controlling the extraction and analysis of debug information generated by the target system T. The host system H and 
target system T of the disclosed embodiment of the Invention communicate via a serial link 110. Most computers are 
equipped with a serial or parallel Interface which can be inexpensively connected to the debug port 1 00 by means of 
a serial connector 108, allowing a variety of computers to function as a host system H. Alternatively, the serial connector 
1 08 could be replaced with higher speed JTAG-to-networi< conversion equipment. Further, the target system T can be 

40 configured to analyze debug/trace Information internally. 

[0020] Referring now to Figure 2. details of an embedded processor device 102 according to the present Invention 
are provided. In addition to the processor core 104, Rgure 2 depicts various elements of an enhanced embodiment of 
the debug port 100 capable of utilizing and controlling the trace cache 200. Many other configurations are possible, 
as will become apparent to those skilled in the art, and the various processor device 102 components described below 

45 are shown for purposes of Illustrating the benefits associated with providing an on-chip trace cache 200. 

[0021] Of significance to the disclosed embodiment of the Invention, the trace control circuitry 21 6 and trace cache 
200 operate to provide trace Infomiatlon for reconstmctlng Instruction execution flow In the processor core 1 04. The 
trace control circuitry 218 supports "tracing" to a trace pad Interface port 220 or to the instruction trace cache 200 and 
provides user control for selectively activating instruction trace capture. Other features enabled by the trace control 

so circuitry 21 8 include programmabllrty of synchronization address generation and user specified trace records, as dis- 
cussed In greater detail below. The trace control circuitry 21 B also controls a trace pad Interface port 220. When utilized, 
the trace pad Interface port 220 Is capable of providing trace data while the processorcore 1 04 Is executing instmctions, 
although clock synchronization and other Issues may arise. The instruction trace cache 200 addresses many of these 
Issues, Improving bandwidth matching and alleviating the need to Incorporate throttle-back circuitry In the processor 

55 core 104. 

[0022] At a minimum, only the conventional JTAG pins need be supported In the software debug port 100 tn the 
described embodiment of the Invention. The JTAG pins essentially become a transportation mechanism, using existing 
pins, to enter commands to be perfomied by the processorcore 104. More specifically, the test clock signal TCK, the 
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test mode select signal TMS, the test data Input signal TDI and the test data output signal TOO provided to and driven 
by the JTAG Test Access Port (TAP) controller 204 are conventional JTAG support signals and known to those skilled 
in the art. As discussed In more detail below, an "enhanced" embodiment of the debug port 100 adds the command 
acknowledge signal CMDACK, the break request/trace capture signal BRTC, the stop transmit signal STOPTX, and 

5 the trigger signal TRIG to the standard JTAG interface. The additional signals allow for pinpoint accuracy of external 
breakpoint assertion and monitoring, triggering of extemal devices in response to Intemal breakpoints, and elimination 
of status polling of the JTAG serial interface. These "sideband" signals offer extra functionality and improve commu- 
nications speeds for the debug port 1 00. These signals also aid in the operation of an optional parallel port 21 4 provided 
on special bond-out versions of the disclosed embedded processor device 102. 

10 [0023] Via the conventional JTAG signals, the JTAG TAP controller 204 accepts standard JTAG serial data and 
control. When a DEBUG Instruction has been written to the JTAG Instmction register, a serial debug shifter 212 is 
connected to the JTAG test data input signal TDI and test data output signal TDO, such that commands and data can 
then be loaded Into and read from debug registers 210. In the disclosed embodiment of the Invention, the debug 
registers 210 Include two debug registers for transmitting (TX.DATA register) and receiving (RX.DATA register) data, 

'5 an Instruction trace-conflgaratlon register (rfCR), and a debug control status register (OCSR), 

[0024] A control Interface state machine 206 coordinates the loading/reading of data to/from the serial debug shifter 
212 and the debug registers 210. A command decode and processing block 208 decodes commands/data and dis- 
patches them to processor Interface logte 202 and trace debug interface logic 216. In addition to perfomiing other 
functions, the trace debug Interface logic 21 6 and trace control logic 218 coordinate the communication of trace Infor- 

20 mation from the trace cache 200 to the TAP controller 204. The processor interface logic 202 communicates directly 
with the processor core 104, as well as the trace control logic 218. As described more fully below, parallel port logic 
214 communicates with a control interface state machine 206 and the debug registers 210 to perform parallel data 
read/write operations in optional bond-out versions of the embedded processor device 102. 
[0025] Before debug infomiatlon is communicated via the debug port 1 00 using only conventional JTAG signals, the 

2S port 1 00 Is enabled by writing the publte JTAG instruction DEBUG Into a JTAG Instruction register contained within the 
TAP controller 204. As shown below, the JTAG instruction register of the disclosed embodiment Is a 38-bit register 
comprising a 32-blt data field (debug_data[31:0]), a four-bit command field to point to various Intemal registers and 
functions provided by the debug port 100, a command pending flag, and a command finished flag. It is possible for 
some commands to use bits from the debug_data field as a sub-field to extend the number of available commands. 

30 
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[0026] This JTAG instruction register is selected by toggling the test mode select signal TMS. The test mode select 
40 signal TMS allows the JTAG path of clocking to be changed in the scan path, enabling multiple paths of varying lengths 
to be used. Preferably, the JTAG instruction register Is accessible via a short path. This register is configured to Include 
a "soft" register for holding values to be loaded into or received from specified system registers. 
[0027] Referring now to Figure 3, a simplified block diagram depicting the relationship between an exemplary In- 
struction trace cache 200 and other components of an embedded processor device 102 according to the present 
^5 Invention is shown. In one contemplated embodiment of the Invention, the trace cache 200 is a 1 28 entry first-ln, first- 
out (FIFO) circular cache that records the most recent trace entries. Increasing the size of the trace cache 200 Increases 
the amount of Instruction trace Information that can be captured, although the amount of required silicon area may 
Increase. 

[0028] As described in more detail below, the trace cache 200 of the disclosed embodiment of the Invention stores 
so a plurality of 20-bit (or more) trace entries Indicative of the order In which Instructions are executed by the processor 
core 104. Other Infonnation, such as task Identifiers and trace capture stop/start Information, can also be placed in the 
trace cache 200. The contents of the trace cache 200 are provided to external hardware, such as the host system H, 
via either serial or parallel trace pins 230. Alternatively, the target system T can be configured to examine the contents 
of the trace cache 200 Internally. 
55 [0029] Figure 4 provides a high-level flow chart of command passing when using a standard JTAG interface. Upon 
entering debug mode In step 400 the DEBUG Instruction Is written to the TAP controller 204 in step 402. Next, step 
404, the 38-blt serial value Is shifted in as a whole, with the command pending flag set and desired data (if applicable, 
othenvise zero) jn the data field. Control proceeds to step 406 where the pending command is loaded/unloaded and 
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the command finished flag checked. Completion of a command typically invoh/es transferring a value between a data 
register and a processor register or memory/10 location. After the command has been completed, the processor 104 
clears the command pending flag and sets the command finished flag, at the same time storing a value In the data 
field If applicable. The entire 38-bit register Is scanned to monitor the command finished and command pending flags, 

s If the pending flag is reset to zero and the finished flag is set to one, the previous command has finished. The status 
of the flags Is captured by the control interface state machine 206. A slave copy of the flags' status Is saved internally 
to determine If the next Instruction should be loaded. The slave copy is maintained due to the possibility of a change 
in flag status between TAP controller 204 states. This allows the processor 104 to determine If the previous Instruction 
has finished before loading the next Instruction. 

10 [0030] If the finished flag Is not set as determined In step 408, control proceeds to step 41 0 and the loadlng/unioading 
of the 38-bit command is repeated. The command finished flag Is also checked. Control then returns to step 408. If 
the finished flag Is set as detemilned In step 408, control returns to step 406 for processing of the next command. 
DEBUG mode is exited via a typical JTAG process. 

[0031] Returning to Figure 2, the aforementioned optional sideband signals are utilized In the enhanced debug port 
15 100 to provide extra functtenallty. The opttonal sfdebamf signals indtide a break requestrtrace capture signal BRTC 
that can function as a break request signal or a trace capture enable signal depending on the status of bit set In the 
debug control/status register. If the break request/trace capture signal BRTC Is set to function as a break request 
signal, It is asserted to cause the processor 104 to enter debug mode (the processor 104 can also be stopped by 
scanning in a halt command via the convention JTAG signals). If set to function as a trace capture enable signal, 
20 asserting the break requestrtrace capture signal BRTC enables trace capture. Deasserting the signal turns trace capture 
off. The signal takes effect on the next Instruction boundary after it Is detected and Is synchronized with the Internal 
processor clock. The break request/trace capture signal BRTC may be asserted at any time. 
[0032] The trigger signal TRIG is configured to pulse whenever an Internal processor breakpoint has been asserted. 
The trigger signal TRIG may be used to trigger an external capturing device such as a logic analyzer, and Is synchro- 
is nized with the trace record capture clock signal TRACECLK. When a breakpoint Is generated, the event Is synchronized 
with the trace capture clock signal TRACECLK, after whteh the trigger signal TRIG Is held active for the duration of 
trace capture. 

[0033] The stop transmit signal STOPTX is asserted when the processor 104 has entered DEBUG mode and is 
ready for register interrogation/modification, memory or I/O reads and writes through the debug port 100, In the dis- 
30 closed embodiment of the Invention, the stop transmit signal STOPTX reflects the state of a bit In the debug control 
status register (DCSR). The stop transmit signal STOPTX is synchronous with the trace capture clock signal TRACE- 
CLK. 

[0O34] The command acknowledge signal CM DACK Is described In conjunction with Figure 5, which shows simplified 
command passing In the enhanced debug port 1 00 of Figure 2. Again, to place the target system T Into DEBUG mode, 

35 a DEBUG instruction Is written to the TAP controller 204 in step 502. Control proceeds to step 504 and the command 
acknowledge signal CMDACK is monitored by the host system H to detennlne command completion status. This signal 
Is asserted high by the target system T simultaneously with the command finished flag and remains high until the next 
shift cycle begins. When using the command acknowledge signal Cf^DACK, it is not necessary to shift out the JTAG 
Instruction register to capture the command finished flag status. The command acknowledge signal CMDACK transl- 

40 tions high on the next rising edge of the test clock signal TCK after the command finished flag has changed from zero 
to one. When using the enhanced JTAG signals, a new shift sequence (step 506) Is not started by the host system H 
until the command acknowledge signal CIVIDACK pin has been asserted high. The command acknowledge signal 
CMDACK Is synchronous with the test clock signal TCK. The test clock signal TCK need not be clocked at all times, 
but is ideally clocked continuously when waiting for a command acknowledge signal CMDACK response. 

45 

OPERATING SYSTEM/APPLICATION COMMUNICATION VIA THE DEBUG PORT 100 

[0035] Also Included In debug register block 210 is an Instruction trace configuration register (ITCR). This 32-bit 
register provides for the enabling/disabling and configuration of Instruction trace debug functions. Numerous such 
so functions are contemplated, Including various levels of tracing, trace synchronization force counts, trace Initialization, 
Instruction tracing modes, clock divider ratio information, as well as additional functions shown in the following table. 
The ITCR is accessed through a JTAG instruction register write/read command as is the case with the other registers 
of the debug register block 210, or via a reserved instruction. 

55 
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Instruction Trace Configuration Register (ITCR). 




BIT 


SYMBOL 


DESCRIPTION/FUNCTION 


5 


31:30 


Reserved 


Reserved 




29 


RXINTEN 


Enables interrupt when RX bit Is set 




28 


TXINTEN 


Enables intemtpt when TX bit is set 


10 


27 


TX 


Indicates that the target system T Is ready to transmit data to the host system H and the 
data Is available in the TX.DATA register 




26 


RX 


indicates that data has been received from the host and placed in the RX.DATA register 




25 


DISL1TR 


Disables level 1 tracing 


15 


24 


OlSLOTR 


Disables level 0 tracing 




23 


DISCSB 


Disables current segment base trace record 




22:16 


TSYNC[6:0] 


Sets the maximum number of Branch Sequence trace records that may be output by the 
u dcs coniroi diock d\ o oeiore a syncnronizing address record is lorced 


20 


15 


TSR3 


oois ut ci9eii9 iicico muae on uno irap 




14 


TSR2 


^Atc Of /^loQrc trar^Q n^r\Ha r\n P^DO trtsF* 

oc;u> ur vmaib irdwa moue on ur\d. irap 




13 


TSR1 


SAts nroioars frraoA mnria nn DQ1 tran 


25 


12 


TSRO 


Sets or cIPflPQ traf^o mnHA nn r^Rn \rar\ 




11 


TRACF3 


ciioijies 1 laco moue loggiiny using urio 




10 


TRACE2 


Enables Trace mode toggling using DR2 




9 


TRACE1 


Enables Trace mode toggling using DR1 


on 
OU 


8 


TRACEO 


Enables Trace mode toggling using DRO 




7 


TRON 


Trace on/off 




6:4 


TCLK[2:0] 


Encoded dh^lder ratio between internal processor clock and TRACECLK 


35 


3 


ITM 


Sets internal or external (bond-out) instruction tracing mode 




2 


TINIT 


Trace initialization 




1 


TRIGEN 


Enables pulsing of external trigger signal TRIG following receipt of any legacy debug 
breakpoint; Independent of the Debug Trap Enable function In the DCSR 


40 


0 


GTEN 


Global enable for Instniction tracing through the internal trace buffer or via the external 
(bond'out) interface 



[0036J Another debug register, the debug control/status register (DCSR), provides an indication of when the proc- 
essor 1 04 has entered debug mode and allows the processor 104 to be forced into DEBUG mode through the enhanced 

45 JTAG Interface. As shown In the following table, the DCSR also enables miscellaneous control features, such as: 
forcing a ready signal to the processor 1 04, controlling memory access space for accesses Initiated through the debug 
port, disabling cache flush on entry to the DEBUG mode, the TX and RX bits, the parallel port 214 enable, forced 
breaks, forced global reset, and other functions. The ordering or presence of the various bits in either the ITCR or 
DCSR is not considered critical to the operation of the Invention. 

50 



Debug Control/Status Register (DCSR). 


BIT 


SYMBOL 


PESCRIPTION/FUNCTION 


31:12 


Reserved 


Reserved 


11 


TX 


Indicates that the target system T is ready to transmit data to the host system H and the 
data Is available In theTX_DATA register 
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(continued) 



Debug Control/Status Register (DCSR). 


BIT 


SYMBOL 


DESCRIPTION/FUNCTION 


10 


RX 


Indicates that data has been received from the host and placed In the RX.DATA register 


9 


DISFLUSH 


Disables cache flush on entrv to DEBUG moda 


8 


SMMSP 


Controls memory access space (nomial memory space/ system management mode 
memorv) for accesses initiated throuah the Dabufi Port loo 


7 


STOP 


indicates whether the orocessor lOd la In r^PRtJA mnria /onuK/alont tA etnn tranemit fttnnal 
STOPTX 


6 


FRCRDY 


Forces the ready signal RDY to the processor 1 04 to be pulsed for one processor clock; 
useful when It js apparent that the processor 104 is stalled waiting for a ready signal from 
a non-responding device 


5 


BRKMODE 


Selects the function of the breaic request/trace capture signal BRTC (brealc request or trace 
capture on/off) 


4 


DBTEN 


Enables entry to debug mode or toggle trace mode enable on a trap/fault via processor 1 04 
registers DR0-DR7 or other legacy debug trap/fault mechanisms 


3 


PARENB 


Enables parallel port 214 


2 


DSPC 


Disables stopping of interna! processor clocks in the Halt and Stop Grant states 


1 


FBRK 


Forces processor 104 into DEBUG mode at the next Instruction boundary (equivalent to 
pulsing the external BRTC pin) 


0 


PRESET 


Forces global reset 



30 l^^^^ When in cross debug environment such as that of Figure 1 , it is necessary for the parent task running on the 
target system T to send information to the host platform H controlling it. This data may consist, for example, of a 
character stream from a printfQ call or register infomaation from a Task's Control Block (TCB). One contemplated 
method f or transf ening the data Is for the operating system to place the data in a known region, then via a trap instruction, 
cause DEBUG mode to be entered. 

[0038] Via debug port 100 commands, the host system H can then determine the reason that DEBUG mode was 
entered, and respond by retrieving the data from the reserved region. However, while the processor 1 04 Is in DEBUG 
mode, nomial processor execution is stopped. As noted above, this is undesirable for many real-time systems. 
[0039] This situation is addressed according to the present invention by providing two debug registers in the debug 
port 1 00 for transmitting (TX.DATA register) and receiving (RX.DATA register) data. These registers can be accessed 
^ using the soft address and JTAG instruction register commands. As noted, after the host system H has written a debug 
instruction to the JTAG instruction register, the serial debug shifter 21 2 Is coupled to the test data input signal TDl line 
and test data output signal TDO line. 

[0040] When the processor 1 04 executes code causing it to transmit data, it first tests a TX bit in the ITCR. If the TX 
bit Is set to zero then the processor 1 04 executes a processor Instnjction (either a memory or I/O write) to transfer the 
data to the TX^DATA register. The debug port 100 sets the TX bit in the DCSR and ITCR, Indicating to the host system 
H that it is ready to transmit data. Also, the STOPTX pin is set high. After the host system H completes reading the 
transmit data from the TX_DATA register, the TX bit Is set to zero. A TXINTEN bit in the ITCR Is then set to generate 
a signal to interrupt the processor 1 04. The Interrupt is generated only when the TX bit In the ITCR transitions to zero. 
When the TXINTEN bit Is not set, the processor 104 polls the ITCR to detennlne the status of the TX bit to further 
transmit data. 

[0041] When the host system H desires to send data, it first tests a RX bit in the ITCR. if the RX bit Is set to zero, 
the host system H writes the data to the RX_DATA register and the RX bit is set to one In both the DCSR and ITCr! 
A RXINT bit is then set In the ITCR to generate a signal to interrupt the processor 1 04. This intenrupt Is only generated 
when the RX in the ITCR transitions to one. When the RXINTEN bit is not set, the processor 104 polls the ITCR to 
verify the status of the RX bit. If the RX bit is set to one, the processor instruction Is executed to read data from the 
RX.DATA register After the data is read by the processor 104 from the RX.DATA register the RX bit is set to zero. 
The host system H continuously reads the ITCR to detemnine the status of the RX bit to further send data. 
[0042] This technique enables an operating system or application to communicate with the host system H without 
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stopping processor 1 04 execution. Communication is convenientiy achieved via the debug port 1 00 with minimal impact 
to on-chip application resources, in some cases It Is necessary to disable system interrupts. This requires that the RX 
and TX bits be examined by the processor 100. In this situation, the communication link Is driven in a polled mode. 

5 PARALLEL INTERFACE TO DEBUG PORT 1 00 

[0043] Some embedded systems require instruction trace to be examined while maintaining I/O and data processing 
operations. Without the use of a multi-tasking operating system, a bond-out version of the embedded processor device 
102 Is preferable to provide the trace data, as examining the trace cache 200 via the debug port 100 requires the 
processor 1 04 to be stopped. 

[0044] in the disclosed embodiment of the invention, a parallel port 214 is also provided In an optional bond-out 
version of the embedded processor device 102 to provide parallel command and data access to the debug port 100. 
This Interface provides a 16-blt data path that is multiplexed with the trace pad Interface port 220. More specifically, 
the parallel port 214 provides a 16-blt wide bi-directional data bus (PDATA[16:0]), a 3-blt address bus (PADR[2:0]), a 
15 paftdteT debug port readAvrftd street ^rghaT (PRW), a trace valid signal TV and an Instructlcm trace record output clock 
TRACECLOCK (TC), Although not shared with the trace pad Interface port 220, a parallel bus request/grant signal pair 
PBREQ/PBGNT (not shown) are also provided. The parallel port 214 is enabled by setting a bit In the DOSR. Serial 
communications via the debug port 100 are not disabled when the parallel port 214 Is enabled. 

20 



22 21 20 


19 


16 




0 


TV 


TC 


PRW 


PADR [2:0] 


PDATA [15:0] 



25 

Boad-Out Pins/Parallei Port 214 Fonnat 

[0045] The parallel port 21 4 Is primarily Intended for fast downloads/uploads to and from target system T memory. 
However, the parallel port 214 may be used for all debug communications with the target system T whenever the 
30 processor 104 Is stopped. The serial debug signals (standard or enhanced) are used for debug access to the target 
system T when the processor 1 04 is executing Instructions. 

[0046] In a similar manner to the JTAG standard, all inputs to the parallel port 214 are sampled on the rising edge 
of the test clock signal TCK, and ail outputs are changed on the failing edge of the test clock signal TCK. I n the disclosed 
embodiment, the parallel port 214 shares pins with the trace pad Interface 220, requiring parallel commands to be 

35 Initiated only while the processor 1 04 Is stopped and the trace pad Interface 220 Is disconnected from the shared bus. 
[0047] The parallel bus request signal PBREQ and parallel bus grant signal PBGNT are provided to expedite multi- 
plexing of the shared bus signals between the trace cache 200 and the parallel port 214. When the host Interface to 
the parallel port 214 determines that the parallel bus request signal PBREQ is asserted, It begins driving the parallel 
port 214 signals and asserts the parallel bus grant signal PBGNT. 

40 [0048] When entering or leaving DEBUG mode with the parallel port 214 enabled, the parailel port 214 Is used for 
the processor state save and restore cycles. The parallel bus request signal PBREQ is asserted Immediately before 
the beginning of a save state sequence penultimate to entry of DEBUG mode. On the last restore state cycle, the 
parallel bus request signal PBREQ is deasserted after latching the write data. The parailel port 214 host interface 
responds to parallel bus request signal PBREQ deassertion by tri-stating Its parallel port drivers and deasseriiing the 

45 parallel bus grant signal PBGNT The parailel port 214 then enables the debug trace port pin drivers, completes the 
last restore state cycle, asserts the command acknowledge signal CMDACK. and retums control of the Interface to 
tracecontrol logic 218. 

[0049] When communicating via the parailel port 214, the address pins PADR[2:0] are used for selection of the field 
of the JTAG Instruction register, which Is mapped to the 16-bit data bus PDATA[15:6] as shown in the following table: 

50 



PADR[2:0] 


Data Selection 


000 


No selection (null operation) 


001 


4-bit command register; command driven on PDATA(3:0] 


010 


High 16-bits of debug^data 


01 1 


Low 1 6-bits of debug^data 
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(continued) 



PADR(2:0] Data Selection 



100-111 



Reserved 



5 



[00501 't Is not necessary to update both halves of the debua_data [31 :0] register If only one of the halves Is being 
used (e.g., on 8-blt I/O cycle data writes). The command pending flag Is automatically set when perfoimlng a write 
operation to the four-bit command register, and is cleared when the command finished flag is asserted. The host system 
10 H can monitor the command acfcnowledge signal CIVIDACK to determine when the finished flag has been asserted 
Use of the parallel port 214 provides full visibility of execution history, without requiring throttling, of the processor core 
104. The trace cache 200. if needed, can be configured for use as a buffer to the parallel port 214 to alleviate any 
bandwidth matching Issues. 

IS OPERATING SYSTEI^ ANO DEBUGGER INTEGRATION 

[0051] In the disclosed embodiment of the Invention, the operation of all debug supporting features, including the 
trace cache 200, can be controlled through the debug port 1 00 or via processor Instojctlons. These processor instruc- 
tions may be from a monitor program, target hosted debugger, or conventional pod-wear. The debug port 100 performs 
20 data moves which are Initiated by serial data port commands rather than processor Instructions. 

[00521 Operation of the processor from conventional pod-space Is very similar to operating In DEBUG mode from a 
monitor program. All debug operations can be controlled via processor Instructions. It makes no difference whether 
these instructions come from pod-space or regular memory. This enables an operating system to be extended to include 
additional debug capabilities. 

23 [0053] Of course, via privileged system calls such a ptraceO. operating systems have long supported debuggers. 

However, the incorporation of an on-chip trace cache 200 now enables an operating system to offer Instruction trace 

capability. The ability to trace Is often considered essential In real-time applications. In a debug environment according 

to the present invention, it Is possible to enhance an operating system to support limited trace without the Incorporation 

of an "external" logic analyzer or In-circult emulator. 
30 [0054] Examples of Instructions used to support internal loading and retrieving of trace cache 200 contents Include 

a load instruction trace cache record command LITCR and a store Instruction trace cache record command SITCR. 

The command LITCR loads an Indexed record In the trace cache 200, as specified by a trace cache pointer ITREC. 

PTR, Witt! the contents of the EAX register of the processor core 104. The trace cache pointer ITREC.PTR is pre- 

Incremented, such that the general operation of the command LITCR is as follows: 



ITREC.PTR <- ITREC.PTR +1 ; 

iTRECf/rflEC.prft;<- eax. 

In ttie event that the Instruction trace record (see description of trace record fonnat below) is smaller that the EAX 
40 record, only a portion of the EAX register is utilized. 

[0055] Similarly, the store Instruction trace cache record command SITCR Is used to retrieve and store (in the EAX 
register) an indexed record from tfie trace cache 200. The contents of the ECX register of the processor core 1 04 are 
used as an offset tiiat Is added to the trace cache pointer ITREC.PTR to create an Index Into the trace cache 200. The 
ECX register is post-Incremented while the trace cache pointer ITREC.PTR is unaffected, such that: 



EAX <- n'HEC[ECX+ tTREC.PTRl 
ECX<-ECX+1. 

Numerous variations to the format of the LITCR and SITCR commands will be evident to those skilled art. 

50 [0056] Extending an operating system to support on-chip trace has certain advantages within the communications 
Industry. It enables the system I/O and communication activity to be maintained while a task is being traced. Tradition- 
ally, the use of an in-circuit emulator has necessitated ttiat ttie processor be stopped before the processor's state and 
trace can be examined [unlike ptraceQ). This disnjpts continuous support of I/O data processing. 
[0057] Additionally, the trace cache 200 is very useful when used with equipment In the field. If an unexpected system 

55 crash occurs, the trace cache 200 can be examined to observe the execution history leading up to the crash event 
When used in portable systems or other environments in which power consumption Is a concern, the trace cache 200 
can be disabled as necessary via power management circuitry. 



35 



45 
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EXEMPLARY TRACE RECORD FORMAT 

[0058] !n the disclosed embodiment of the Invention, an Instruction trace record Is 20 bits wide and consists of two 
fields, TCODE (Trace Code) and TDATA (Trace Data), as well as a valid bit V. The TCODE field is a code that Identifies 
5 the type of data in the TDATA field. The TDATA field contains software trace information used for debug purposes. 





20 I 


9 15 




0 


10 


V 


TCODE (Trace Code) 


TDATA (Trace Data) 



Instruction Trace Record Format. 



IS [0059] In one contemjDtated Embodiment of the invention, the embedded processor device 1 02 reports eleven dif- 
ferent trace codes as set forth In the following table: 



TCODE# 


TCODE Type 


TDATA 


0000 


Missed Trace 


Not Valid 


0001 


Conditional Branch 


Contains Branch Sequence 


0010 


Branch Target 


Contains Branch Target Address 


0O11 


Previous Segment Base 


Contains Previous Segment Base Address and Attributes 


0100 


Curent Segment Base 


Contains Current Segment Base Address and Attributes 


0101 


Intenupt 


Contains Vector Number of Exception or Intenupt 


0110 


Trace Synchronization 


Contains Address of Most Recently Executed Instruction 


0111 


Multiple Trace 


Contains 2nd or 3rd Record of Entry With Multiple Records 


1000 


Trace Stop 


Contains Instruction Address Where Trace Capture Was Stopped 


1001 


User Trace 


Contains User Specified Trace Data 


1010 


Perfomiance Profile 


Contains Perfomiance Profiling Data . 



[0060] The trace cache 200 is of limited storage capacity; thus a certain amount of "compression" in captured trace 
data Is desirable, in capturing trace data, the following discussion assumes that an image of the program being traced 
is available to the host system H. if an address can be obtained from a program image (Object Module), then it Is not 
provided In the trace data. Preferably, only Instructions which disrupt the instoiction flow are reported; and further, only 
those where the target address Is in some way data dependent. For example, such "dismpting" events Include call 
instructions or unconditional branch instructions In which the target address is provided from a data register or other 
memory location such as a stacl<. 

[0061] As Indicated in the preceding table, other desired trace Infonnatlon includes: the target address of a trap or 
intemjpt handler; the target address of a return instoiction; a conditional branch instruction having a target address 
which Is data register dependent (otherwise, all that is needed Is a 1 -bit trace indicating if the branch was taken or not) ; 
and, most frequently, addresses from procedure returns. Other Infomnation, such as task identifiers and trace capture 
stop/start information, can also be placed in the trace cache 200. The precise contents and nature of the trace records 
are not considered critical to the Invention.- 

[0062] Figure 6A illustrates an exemplary fomfiat for reporting conditional branch events. In the disclosed embodiment 
of the invention, the outcome of up to 15 branch events can be grouped Into a single trace entry. The 16-bit TDATA 

field (or "BFIELD") contains 1 -bit branch outcome trace entries, and Is labeled as a TCODE = 0001 entry. The TDATA 

field Is Initially cleared except for the left most bit, which Is set to 1 . As each nevy conditional branch is encountered, a 

new one bit entry is added on the left and any other entries are shifted to the right by one bit. 

[0063] Using a 1 28 entry trace cache 200 allows 320 bytes of infonnation to be stored. Assuming a branch frequency 

of one branch every six instructions, the disclosed trace cache 200 therefore provides an effective trace record of 1 ,536 

instructions. This estimate does not take Into account the occurrence of call, jump and return instructions. 

[0064] In the disclosed embodiment of the invention, the trace control logic 218 monitors instruction execution via 
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processor Interface logic 202. When a branch target address must be reported, Information contained within a current 
conditional branch TDATA field Is mar1<ed as complete by the trace control logic 218, even if 15 entries have not 
accumulated. As shown In Figure 68, the target address (In a processor-based device 102 using 32-bit addressing) Is 
then recorded In a trace entry pair, with the first entry (TCODE = 001 0) providing the high 1 6-bits of the target address 
5 and the second entry (TCODE = 0111) providing the low 16-bits of the target address. When a branch target address 
is provided for a conditional jump instruction, no 1 -bit branch outcome trace entry appears for the reported branch. 

STARTING AND STOPPING TRACE CAPTURE 

10 [0065] Referring now to Rgure 6C, It may be desirable to start and stop trace gathering during certain sections of 
program execution; for example, when a tasic context switch occurs. When trace capture is stopped, no trace entries 
are entered Into the trace cache 200, nor do any appear on the bond-out pins of trace port 21 4. Different methods are 
contemplated for enabling and disabling trace capture. For example, an x86 command can be provided, or an existing 
x86 command can be utilized to toggle a bit In an I/O port location. Altemath^ely, on-chip breai^polnt control registers 

is (not showrt)«anbe configared to indicate the-addressesAAfheretrace captureshoufd startfetop-. Wherrtraclngls halted, 
a trace entry (TCODE = 1000, TCODE = 0111) recording the last trace address is placed in the trace stream. When 
tracing is resumed, a trace synchronization entry (TCODE = 0110, TCODE = 0111) containing the address of the 
currently executing Instruction Is generated. 

[0066] it may be important to account for segment changes that occur while tracing is stopped. This situation can 
20 be partially resolved by selecting an option to immediately follow a TCODE = 1 000 entry with a cunrent segment base 
address entry (TCODE = 0100, TCODE = 0111), as shown In Figure 6C. A configuration option Is also desirable to 
enable a current segment base address entry at the end of a trace prior to entering Debug mode. By contrast, it may 
not be desirable to provide segment base infonnation when the base has not changed, such as when an intenrupt has 
occun-ed. 

25 [0067] Referring to Figure 6D, following the occunrence of an asynchronous or synchronous event such as an Intenrupt 
or trap, a TCODE = 01 01 trace entry is generated to provide the address of the target Intermpt handler However, It Is 
also desirable to record the address of the Instruction which was Intenrupted by generating a trace synchronization 
(TCODE = 0110) entry Immediately prior to the intenupt entry, as well as the previous segment base address (TCODE 
= 001 1 ). The trace synchronization entry contains the address of the last Instruction retired before the Interrupt handler 

30 commences. 

SEGIVIENT CHANGES 

[0068] Figure 6E Illustrates a trace entry used to report a change In segment parameters. When processing a trace 
35 stream In accordance with the invention, trace address values are combined with a segment base address to detemilne 
an instruction's linear address. The base address, as well as the default data operand size (32 or 16-bit mode), are 
subject to change. As a result, the TCODE = 001 1 and 01 1 1 ientries are configured to provide the Infonnation necessary 
to accurately reconstruct Instruction flow. The TDATA field corresponding to a TCODE = 0011 entry contains the high 
1 6-bits of the previous segment base address, while the associated TCODE = 0111 entry contains the low 1 5 or 4 bits 
40 (depending on whether the instruction Is executed In real or protected mode). The TCODE =0111 entry also preferably 
Includes bits indicating the current segment size (32-bit or 16-blt), the operating mode (real or protected), and a bit 
Indicating whetherpaging Is being utilized. Segment Infonmatlon generally relates tothe previous segment, notacurrent 
(target) segment. Cunrent segment Infonnation Is obtained by stopping and examining the state of the processor core 
104. 

45 

USER SPECIFIED TRACE ENTRY 

[0069] There are circumstance when an application program or operating system may wish to add additional infor- 
mation into a trace stream. For this to occur, an x86 instruction is preferably provided which enables a 1 6-bit data value 

so to be placed in the trace stream at a desired execution position. The Instruction can be implemented as a move to 1/ 
0 space, with the operand being provided by memory or a register. When the processor core 104 executes this In- 
struction, the user specified trace entry is captured by the trace control logic 218 and placed In the trace cache 200. 
As shown In Figure 6F, a TCODE = 1 001 entry is used for this purpose in the disclosed embodiment of the Invention. 
This entry might provide, for example, a previous or current tasic Identifier when a tasi< switch occurs in a multi-tasl^ing 

ss operating system. 
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SYNCHRONIZATION OF TRACE DATA 

[0070] When executing typical software on a processor-based device 102 according to the disclosed embodiment 
of the Invention, few trace entries contain address values. Most entries are of the TCODE = 0001 format, In which a 
5 single bit Indicates the result of a conditional operation. When examining a trace stream, however, data can only be 
studied in relation to a known program address. For example, starting with the oldest entry in the trace cache 200, ail 
entries until an address entry are of little use. Algorithm synchronization typically begins from a trace entry providing 
a target address. 

[0071] If the trace cache 200 contains no entries providing an address, then trace analysis cannot occur. This situation 
10 Is rare, but possible. For this reason, a synchronization register TSYNC Is provided In the prefenBd embodiment of 
invention to control the injection of synchronizing address information. If the synchronization register TSYNC Is set to 
zero, then trace synchronization entries are not generated. 



15 



20 



45 



TSYNC (Trace Synchronization) 



Trace Entry Synchronization Entry Control Register. 



[0072] Figure 6G depicts an exemplary trace synchronization entfy. In operation, a counter register is set to the value 
contained in the synchronization register TSYNC whenever a trace entry containing a target address is generated. 

25 The counter is decremented by one for all other trace entries. If the counter reaches zero, a trace entry is Inserted 
(TCODE = 01 1 0) containing the address of the most recently retired instruction (or, alternatively, the pending instruc- 
tion). In addition, when a synchronizing entry is recorded In the trace cache 200, it also appears on the trace pins 220 
to ensure sufficient availability of synchronizing trace data for full-function ICE equipment. 
[0073] Trace entry information can also be expanded to include data relating to code coverage or execution perfomi- 

30 ance. This Information is useful, for example, for code testing and perfomiance tuning. Even without these enhance- 
ments, it Is desirable to enable the processor core 104 to access the trace cache 200. In the case of a microcontroller 
device, this feature can be accomplished by mapping the trace cache 20O within a portion of I/O or memory space. A 
more general approach Involves including an instaictlon which supports moving trace cache 200 data into system 
memory. 

35 [0074] Thus, a processor-based device providing a flexible, hlgh-perfomnance solution for furnishing Instruction trace 
information has been described. The processor-based device incorporates an instruction trace cache capable of pro- 
viding trace Infomiation for reconstructing Instruction execution flow on the processor without halting processor oper- 
ation. Both serial and parallel communication channels are provided for communicating trace data to external devices. 
The disclosed on-chip instruction trace cache alleviates various of the bandwidth and clock synchronization problems 

40 that arise In many existing solutions, and also allows less expensive external capture hardware to be utilized. 

[0075] The foregoing disclosure and description of the Invention are illustrative and explanatory thereof, and various 
changes In the size, shape, materials, components, circuit elements, wiring connections and contacts, as well as In 
the details of the illustrated circuitry and constmction and method of operation may be made without departing from 
the spirit of the invention. 



Claims 



1. An electronic processor-based device (102) adapted to execute a series of Instructions obtained from external 
50 sources (1 06), the processor-based device being provided with pins to pennit connection to extemal conductors, 
the electronic processor-based device being characterized by: 

a trace cache (200) coupled to a processor core (104) for storing trace infonnation Indicative of the order In 
which the instructions are executed by the processor core, the trace cache comprising a series of storage 
ss elements, each storage element being adapted to store trace Infonnation, the trace infonnation Including a 

plurality of instruction trace records containing address and data Information, the trace cache (200) being 
configured to load data from the processor core (104) In response to a load command and configured for the 
processor core (1 04) to retrieve data from the trace cache (200) in response to a retrieve command; 
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and a communication channel connected between the trace cache (200) and selected ones of the pins to 
provide for transmission of trace information from the trace cache to extennal devices. 

2. The processor-based device of ciaim 1 wherein the instruction trace records each include a trace data field (TDA- 
5 TA), and a trace code field (TCODE) for storing a code to identify the type of data In the trace data field. 

3. The processor-based device of claim 1 or 2 configured to omit from stored trace Infomiation instruction trace 
records for Instructions between Instructions that disrupt the Instruction flow. 

10 4. The processor-based device of ciaim 1 , 2 or 3 wherein information concern ing executed branch instructions having 
a target address that Is not data register dependent is stored in the trace cache In the fomfi of a single bit. 

5. The processor-based device of any preceding claim, wherein the trace cache (200) Is further configured to provide 
trace capture start/stop Information. 

13 

6. The processor-based device of any preceding claim, wherein the trace cache (200) is further configured to peri- 
odically capture a synchronizing entry, the synchronizing entry being the address of the instruction most recently 
executed by the processor core (1 04). 

20 7. The processor-based device of any preceding claim, wherein the trace cache (200) Is further configured to provide 
intenupt or exception vector infonmatlon. 

8. The processor-based device of any preceding claim, wherein each Instruction trace record further Includes a data 
valid bit (V). 

25 

9. The processor-based device of any preceding claim, wherein the trace cache (200) Is further configured to provide 
task identifier information. 

10. The processor-based device of any preceding claim, wherein the contents of the trace cache (200) are retrievable 
30 by an operating system 

11. The processor-based device of any preceding claim, wherein the communication Interface comprises a serial in- 
terface (204) which Is essentially compliant with the IEEE-1149.1-1990 JTAG interface standard or other similar 
standard. 

35 

12. The processor-based device of any preceding claim, wherein the trace cache (200) Is a first-in, first-out (FIFO) 
circular cache. 

13. The processor-based device of any of claims 1 to 12, further comprising: 

40 

a processor interface (202) coupled to the processor core (1 04) and the trace cache, wherein the trace cache 
Is adapted to load the trace information from the processor core (104) via the processor Interface (202), and 
wherein the processor core (104) is adapted to retrieve the trace infonmatlon from the trace cache via the 
processor interface (202). 

45 

14. A method for analysing trace infomiation in a processor-based device (1 02) having a processor core (1 04), com- 
prising the steps of: 

providing a trace cache (200) within the processor-based device (102), the trace cache comprising a series 
50 of storage elements adapted to store trace infonmatlon; 

capturing trace infomnation from the processor core (104) that is indicative of the order in which the series of 
instructions is executed by the processor core; 

storing the trace Information in the trace cache storage elements as instruction trace records; 
retrieving the trace infonmation by the processor core (1 04) from the trace cache storage elements in response 
55 to a retrieve command; and 

loading other infonmatlon from the processor core (1 04) into the trace cache storage elements In response to 
a load command; 

providing a communication channel (230) from the trace cache to selected pins of the processor-based device 
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(102); and 

communicating the trace infomnation from the trace cache (200) to the selected pins via the communication 
channel (230). 

5 15. The method of claim 1 4. wherein the loading and retrieving steps transfer the trace information between the proc- 
essor core (104) and the trace cache via a processor interface (202). 
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